Phase 3: Machine Learning Lesson 5 of 6

Overfitting, Underfitting
and Generalisation

A model that scores 99% on training data and 60% on real data is useless. A model that scores 65% on both is genuinely valuable. This lesson is about understanding why that gap exists and how to close it.

You will learn
What generalisation means and why it is the real goal
The difference between underfitting and overfitting
The bias-variance tradeoff explained plainly
Techniques to prevent overfitting
How cross-validation gives a more honest evaluation

The real goal is never accuracy on training data

When you train a machine learning model, you optimise its performance on the data you showed it. But that is not actually what you care about. You care about what the model does on data it has never seen. Will it work on tomorrow's emails? On next month's customer transactions? On medical images from a different hospital?

That ability to perform well on new, unseen data is called generalisation. It is the actual goal. Training accuracy is just a proxy metric, and a misleading one if you are not careful.

Analogy

Imagine a student preparing for an exam by memorising the exact questions from last year's past papers, word for word. They ace the practice papers. But the real exam has different questions covering the same material. The memoriser fails while the student who actually understood the concepts does fine. Overfitting is the machine learning equivalent of that kind of memorising.

This is why you always split your data into a training set and a test set before you start. The test set is data the model never sees during training. It is your honest proxy for how the model will perform in the real world. You train on the training set. You evaluate on the test set. You never use the test set to make decisions about your model. The moment you do, it is no longer an honest measure of generalisation.

Underfitting and overfitting

Most problems in machine learning sit on a spectrum between two failure modes. On one end, the model is too simple. On the other, it is too complex.

Underfitting
Too simple

Model is too simple to capture the real pattern. It misses obvious structure in the training data and performs poorly on both training and test data.

High bias · Low variance
Good fit
Just right

Model captures the true underlying pattern without getting distracted by noise. Works well on training data and generalises to new data.

Balanced bias · Balanced variance
Overfitting
Memorising noise

Model is too complex and follows the training data perfectly, including its noise and random quirks. Fails badly on new data it hasn't seen before.

Low bias · High variance

Underfitting is usually obvious and easy to fix: your model is too simple, so you make it more complex. Add more features, use a deeper tree, try a more powerful algorithm. The challenge is that fixing underfitting tends to move you toward overfitting, and vice versa. Finding the right point in between is what model development is really about.

The bias-variance tradeoff

Every model has two sources of error: bias and variance. These two tend to pull in opposite directions, and understanding them helps you diagnose problems and choose solutions.

Bias is systematic error from wrong assumptions. A linear model trying to fit a curved pattern will always be off, no matter how much data you give it. It is biased toward straight lines. More data won't fix this; you need a more flexible model.

Variance is sensitivity to the specific training data. A highly flexible model that follows every wiggle in the training data will produce very different results if you retrain it on a slightly different sample. It is unstable. It has memorised quirks that don't generalise.

Bias-variance tradeoff: error as model complexity grows
Model complexity Error Low High Sweet spot High bias High variance Bias Variance Total error

As model complexity increases, bias (blue) decreases but variance (red) increases. Total error (black) is minimised at the sweet spot where neither dominates. The goal is to find a model complex enough to capture the pattern, but not so complex that it memorises the noise.

How to fix overfitting

Overfitting is by far the more common problem in practice, because modern models are powerful enough to memorise any finite dataset if you let them. Here are the main weapons against it.

Fixes overfitting
Get more training data
More data makes it harder to memorise. With 10,000 diverse examples, the model is forced to find genuine patterns. This is the most reliable fix when it is available.
Fixes overfitting
Reduce model complexity
Limit tree depth, reduce the number of features, use fewer layers. Give the model less room to memorise noise. Start simple and add complexity only when you have evidence it helps.
Fixes overfitting
Regularisation
Techniques like L1 and L2 regularisation penalise large parameter values during training, forcing the model to spread its "confidence" and preventing it from assigning huge weights to noise features.
Fixes overfitting
Dropout (neural networks)
During training, randomly switch off a fraction of neurons. This prevents co-adaptation between neurons and forces the network to learn redundant representations that generalise better.
Fixes underfitting
Increase model complexity
Add more features, increase tree depth, add layers to a neural network, or switch to a more powerful algorithm class. The model simply doesn't have enough capacity to capture the pattern.
Fixes underfitting
Feature engineering
Create new, more informative features from existing ones. Add polynomial terms, interaction terms, or domain-specific transformations. Better features can compensate for a simpler model.

Cross-validation: a more honest evaluation

When your dataset is small, a single train-test split can be noisy. If you happen to put all the easy examples in the test set, your accuracy looks great but means nothing. Cross-validation solves this by rotating which data gets used for testing.

In k-fold cross-validation, you split the data into k equal parts (folds). You train on k-1 folds and test on the remaining one. You repeat this k times, each time using a different fold as the test set. You average the results across all k rounds. This gives you a much more reliable estimate of generalisation performance with no data wasted.

Python cross_validation.py
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_iris(return_X_y=True)

# Compare three models: shallow, medium, and deep trees
for depth in [1, 4, 20]:
    model = DecisionTreeClassifier(max_depth=depth, random_state=42)
    scores = cross_val_score(model, X, y, cv=5)
    print(
        f"Depth {depth:<3} | "
        f"CV accuracy: {scores.mean():.3f} (+/- {scores.std():.3f})"
    )
Output
Depth 1 | CV accuracy: 0.673 (+/- 0.036)  ← underfitting
Depth 4 | CV accuracy: 0.953 (+/- 0.022)  ← good fit
Depth 20 | CV accuracy: 0.940 (+/- 0.038)  ← slightly overfitting

Notice that depth 20 is actually slightly worse than depth 4 on average, and more variable (higher standard deviation). The model at depth 20 is overfitting. Depth 4 finds the sweet spot for this dataset. Cross-validation makes this visible in a way that a single train-test split often cannot.

The golden rule of evaluation

Never touch your test set until you have finished making all decisions about your model. Every time you check test performance and then change something, the test set leaks information into your development process. It is no longer an honest measure of generalisation. If you need to tune hyperparameters, use a separate validation set or cross-validation on the training data only. The test set is a one-shot, final-answer measure.

"The most important skill in machine learning is resisting the urge to keep checking your test score."

A principle learned repeatedly, the hard way
Hands-on activity

Watch overfitting happen in real time

You are going to deliberately overfit a model and then observe the training accuracy and validation accuracy diverge as the model memorises the training data. Then you will apply regularisation and watch the gap close.

01 Open the Lesson 3.5 Colab notebook. Load the breast cancer dataset from Scikit-learn and split it 80/20.
02 Train a DecisionTreeClassifier for max_depth values from 1 to 25. Record both train accuracy and test accuracy for each depth. Plot both curves on the same graph.
03 Identify the point where training accuracy keeps rising but test accuracy starts to fall. That is the onset of overfitting. Which depth is optimal?
04 Run 5-fold cross-validation for each depth instead of a single split. Does the cross-validated optimal depth match what you found with the single split?
05 Try RandomForestClassifier at the same complexity levels. Random forests use ensembling to reduce variance. Does the training-test gap look different compared to a single tree?
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Practice Notebook
Run this lesson's code live in Google Colab
All examples + challenge exercises · Free GPU included · No setup required
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.